Observability
Observability is the ability to answer questions about a running system that you did not anticipate when you built it. Monitoring tells you the interoperability layer is returning errors; observability tells you it is returning errors only for one facility's laboratory messages, and only since yesterday's terminology release.
In a health exchange, the operational question is usually which participant is broken right now — and the exchange is the only component positioned to know.
The three signals
| Signal | Answers | Cost |
|---|---|---|
| Logs | What happened in this specific case | High volume, high storage |
| Metrics | How much, how often, how fast — aggregated | Cheap, low cardinality only |
| Traces | Where the time went across services | Moderate; sampling usually required |
A fourth, health-specific: audit records —
who accessed which patient's data. These are not observability data. They
contain personal health information, have legal retention requirements, and must
be stored separately with their own access controls. Do not put AuditEvent
data in the same store as application logs.
The cardinal rule
No personal health data in telemetry.
Logs, metrics and traces are typically shipped to a monitoring system, retained loosely, and read by engineers with no clinical relationship to any patient. A log line containing a patient name or a diagnosis is a disclosure.
Practical rules:
- Log identifiers, never names, addresses or clinical content
- Prefer internal surrogate identifiers to national identifiers in logs
- Never log request or response bodies for clinical payloads by default — and where an integration genuinely needs it for debugging, gate it behind an explicit, time-limited, audited setting
- Never put a patient identifier in a metric label — it is both a privacy problem and a cardinality explosion
- Scrub tokens, credentials and authorisation headers
- Test for leakage: grep the log store for known synthetic patient names as a routine check
The OpenHIM transaction log is a deliberate exception: it persists bodies by design, which is why it must be treated as clinical data storage rather than as logging.
What to instrument in a health exchange
Per integration channel
- Message volume, by source system and message type
- Error rate, by error class — validation failure, identity unresolved, downstream timeout, authorisation denied
- Latency distribution, at p50/p95/p99 — averages hide the cases that matter
- Queue depth and queue age. Depth alone misses the slow drain; a queue with fifty items that are six hours old is a data-freshness incident
- Retry and dead-letter counts
Per participant system
- Last successful message received. This is the single most useful signal in an exchange: a facility that stopped sending three days ago is invisible in aggregate error rates
- Availability of the system's endpoint
- Conformance failure rate, which spikes after either side upgrades
Business-level signals
Technical health is not the same as the exchange working. Instrument outcomes:
- Records successfully linked to a client registry identity, versus routed to the review queue
- Review queue depth and age — see MPI
- Terminology translation failures, by code and source
- Reporting completeness by facility for the current period
- Break-glass invocations
These are the metrics that tell you whether the architecture is delivering value. A dashboard showing 99.9% uptime and a six-week-old review queue is describing a failed exchange.
SLIs, SLOs and SLAs
- SLI — a measured indicator: "proportion of
$matchrequests completing under 500 ms" - SLO — the internal target: "99% over 30 days"
- SLA — the contractual commitment, with consequences; usually looser than the SLO
Set SLOs from clinical consequence, as with availability tiers. Patient lookup during registration has a tight latency SLO because a queue forms at the desk. Overnight HMIS submission does not.
Error budgets work well in this setting: if the SLO is 99.9% over 30 days, about 43 minutes of failure is acceptable, and that budget is what pays for upgrades and change. When it is exhausted, change stops until reliability is restored. This gives a health ministry a defensible, non-arbitrary rule for slowing down deployments.
Health checks
Three kinds, and conflating them causes outages:
- Liveness — is the process alive? If not, restart it.
- Readiness — can it serve traffic now? Remove from the load balancer if not.
- Dependency health — can it reach the registry, the terminology service, the database?
The failure to avoid: a readiness check that fails when a downstream dependency is unavailable. The orchestrator then removes every instance from service, and a degraded dependency becomes a total outage. Report dependency health as a distinct signal; keep readiness about the service itself.
Tooling
| Tool | Role | Licence |
|---|---|---|
| OpenTelemetry | Vendor-neutral instrumentation for traces, metrics and logs — the right default, because it decouples instrumentation from backend choice | Apache 2.0 |
| Prometheus | Metrics collection and alerting | Apache 2.0 |
| Grafana | Dashboards across metrics, logs and traces | AGPL |
| Loki | Log aggregation, indexed by label | AGPL |
| OpenSearch / Elasticsearch | Log search and analytics | Apache 2.0 / SSPL |
| Jaeger / Tempo | Distributed tracing backends | Apache 2.0 / AGPL |
| Alertmanager | Routing, grouping and silencing alerts | Apache 2.0 |
| Uptime Kuma / Blackbox exporter | External endpoint checks | MIT / Apache 2.0 |
Instrument with OpenTelemetry regardless of backend choice. It is the decision that is cheapest to make now and most expensive to retrofit.
Alerting discipline
An alert that does not require a human to act on it immediately should not page anyone. Health operations teams are small; alert fatigue is the reason real alerts get missed.
For each alert define: what is broken, what the clinical or data impact is, what the responder should do first, and how to confirm resolution. If any of those cannot be written, the alert is a dashboard item.
Alert on symptoms, not causes: "laboratory results have not reached the SHR for 30 minutes" is actionable; "CPU is at 80%" is not.
References
- OpenTelemetry — https://opentelemetry.io/
- Prometheus — https://prometheus.io/
- Grafana — https://grafana.com/oss/
- Google SRE Book, on SLOs and error budgets — https://sre.google/books/
- FHIR
AuditEvent— https://hl7.org/fhir/auditevent.html